Multimodal Large Language Models have demonstrated impressive video understanding, yet their ability to reason over long-form narratives is often masked by visual-centric evaluations and inefficient context processing. Existing benchmarks over-rely on visual heuristics while marginalizing auditory cues, effectively red...
Zhao-Yang Wei, Zipeng Wang, Yu-She Cao et al.· 0 citations
Recent advances in world models and video generation have given rise to an emerging reasoning paradigm that leverages video generative models to simulate, predict, and reason about real-world dynamics. We redefine this paradigm as Thinking in Video, where video is not merely an output artifact but a medium for construc...
Experiments on perception-intensive visual reasoning benchmarks show that LUT outperforms previous latent reasoning methods and remains competitive with latent-text interleaved methods with lower annotation cost.
J. Kang, Siyue Chen, Ming-Da Li et al.· 1 citation
Dynamic Evidence-Guided Preference Optimization (DEPO) is proposed, a new framework that enables evidence-aware and adaptive preference learning for Med-LVLMs and introduces Multi-Modal Evidence Perturbation (MEP) to suppress non-causal textual and visual shortcuts and Dispre-ferred Evidence Resampling (DER) to continu...
Zixuan Huang, Zhihong Zhu, Xiaolong Liu et al.· Annual Meeting of the Associ...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.