Spatial reasoning benchmarks evaluate vision-language models across diverse tasks, but task-level scores do not reveal which underlying capabilities account for success or failure. Each task requires recovering spatial evidence, representing geometry, and reasoning over it. We disentangle these capabilities by comparin...
Peng-Zhan Sun, Jun-Bin Xiao, Ramanathan Rajaraman et al.· 0 citations
This paper studies \textbf{thinking--answer consistency} in vision-language models. We focus on Visual Intention Grounding, where a model infers a target object based on a human intention query and predicts a bounding box. We reveal that previous IoU-based reinforcement learning (RL) frameworks suffer from ``thinking d...
Peng-Zhan Sun, Shiu-hong Kao, Shi-Jie Li et al.· 1 citation
Verifier-free reinforcement learning with probability-based rewards offers a promising way to train LLMs on general reasoning tasks where external verifiers are unavailable. Yet the reliability of these rewards, especially in long-horizon reasoning, remains underexplored. This work identifies a length-dependent failure...
Shiu-hong Kao, Yu-Bo Zhao, Zhen Tian et al.· 0 citations
This work proposes SmartRes, a framework that performs efficiency optimization in the pixel space via dynamic resolution routing and introduces a margin-regularized routing objective that increases foreground-background logit separation and improves foreground recall.
H. Sun, Wang-Bo Zhao, Fanyue Wei et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.