Video large language models (VideoLLMs) achieve strong video understanding performance, but their inference remains expensive due to the large number of redundant spatio-temporal visual tokens in long videos. Existing token pruning methods alleviate this cost by reducing redundant tokens, yet most of them rely on segme...
Meng-Jie Zhang, Qi-Hui Zhu, Tao Zhang et al.· 0 citations
It is demonstrated that resolution differences can serve as a simple and scalable source of privileged information, providing an effective and efficient approach to on-policy self-distillation for multimodal large language models.
This work finds that the deep-layer features of a lightweight speculative model exhibit strong consistency with the target model in the selection of critical tokens for recomputation, and proposes SpecCache, which employs deep-layer hidden-state norms from a speculative model as a proxy to guide the critical token sele...
Zijian Wen, Tao Zhang, Shuangwu Chen et al.· Annual Meeting of the Associ...· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.