STRAND is introduced, a benchmark of human-verified object-centric facts that evaluates intermediate reasoning by decomposing queries into sub-questions, distinguishing genuine temporal understanding from coincidental correctness and an object-centric framework that explicitly constructs and reasons over structured obj...
T. Nguyen, Tri Cao, Khoi M. Le et al.· 0 citations
DynaPix (Dynamic Pixels), a benchmark that makes prediction checkable, exposes a temporal-anchoring gap: models attach a prediction to an event far better than to time itself.
T. Nguyen, Vinh-Hien Do, Quynh T. N. Vo et al.· 0 citations
This work proposes Depth-Ordinal Prompting (DOP), a training-free method that converts monocular depth into a single question-targeted ordinal text cue at the queried objects, without adding a depth image, training a module, injecting features, or using labels.
Quynh T. N. Vo, Phuc T. Dao, Cong-Duy Nguyen et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.