STRAND is introduced, a benchmark of human-verified object-centric facts that evaluates intermediate reasoning by decomposing queries into sub-questions, distinguishing genuine temporal understanding from coincidental correctness and an object-centric framework that explicitly constructs and reasons over structured obj...
T. Nguyen, Tri Cao, Khoi M. Le et al.· 0 citations
DynaPix (Dynamic Pixels), a benchmark that makes prediction checkable, exposes a temporal-anchoring gap: models attach a prediction to an event far better than to time itself.
T. Nguyen, Vinh-Hien Do, Quynh T. N. Vo et al.· 0 citations
This work proposes Depth-Ordinal Prompting (DOP), a training-free method that converts monocular depth into a single question-targeted ordinal text cue at the queried objects, without adding a depth image, training a module, injecting features, or using labels.
Quynh T. N. Vo, Phuc T. Dao, Cong-Duy Nguyen et al.· arXiv.org· 0 citations
Comprehensive in-domain and out-of-domain evaluations across multiple LLMs demonstrate that SFS-DPO and SFS-DPO-R consistently outperform prior step-level training baselines and reveal improvements in self-correction frequency and effectiveness.
Duc Anh Vu, Nhat M. Hoang, Do-Xuan Long et al.· 1 citation
Curvature-Conditioned Query modifies only the read step and is composable with any linear-attention backbone, and improves perplexity, zero-shot downstream accuracy, S-NIAH retrieval at and beyond the training context, length-extrapolation perplexity from 4K to 20K, and LongBench accuracy.
D. Le, Thong Nguyen, Cong-Duy Nguyen et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.