Streaming video understanding requires models to process unbounded visual streams while preserving rich visual semantics across vast temporal horizons, posing a fundamental challenge for memory modeling. Existing approaches primarily focus on increasing memory capacity, either by compressing historical information into...
Yin-Ying Li, Yu-Qian Fu, Yu-Lin Dai et al.· 0 citations
Although vision-language models (VLMs) possess strong visual understanding and reasoning capabilities, existing vision-and-language navigation (VLN) agents struggle to connect semantic reasoning with spatial execution. Two coupled gaps remain in this connection, as intermediate reasoning is not explicitly anchored to v...
Kailing Li, Yu Han, Tian-Wen Qian et al.· 0 citations
This work proposes Change and Invariance Motion Editing (CIME), a unified framework that comprehensively decouples change and invariance into spatial pose and temporal rhythm dimensions and introduces the Riemannian Non-uniform Integral Manifold Mapping module.
Shaohui Lin, Zhenwu Shi, Jingyu Gong et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.