Spatial reasoning benchmarks evaluate vision-language models across diverse tasks, but task-level scores do not reveal which underlying capabilities account for success or failure. Each task requires recovering spatial evidence, representing geometry, and reasoning over it. We disentangle these capabilities by comparin...
Peng-Zhan Sun, Jun-Bin Xiao, Ramanathan Rajaraman et al.· 0 citations
This paper studies \textbf{thinking--answer consistency} in vision-language models. We focus on Visual Intention Grounding, where a model infers a target object based on a human intention query and predicts a bounding box. We reveal that previous IoU-based reinforcement learning (RL) frameworks suffer from ``thinking d...
Peng-Zhan Sun, Shiu-hong Kao, Shi-Jie Li et al.· 1 citation
Experiments show that MyBuddy significantly enhances the performance of foundation models on BuddyVQA, a benchmark for companion-style question answering (QA) on egocentric streaming video, and generalize to other streaming and common video QA benchmarks, demonstrating the applicability and effectiveness of the approac...
Hangyu Qin, Jun-Bin Xiao, Sheng Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.