Frame selection is essential for applying Large Multimodal Models (LMMs) to long videos due to severe frame redundancy and limited context windows. Since the appropriate frame budget varies with the downstream LMM, reasoning demands, and latency constraints, a practical selector should serve multiple budgets. However,...
Wang Chen, Yu Chen, Xiang Wang et al.· 0 citations
WaveZip is proposed, a joint signal-frequency-domain framework for efficient video inference that requires no task-specific training and can be seamlessly integrated into off-the-shelf LVLMs to boost inference efficiency.
Yuhui Zeng, Wang Chen, Jin-Fa Huang et al.· arXiv.org· 1 citation
TimePLE is proposed, which reformulates VTG from endpoint prediction to interval-native grounding by predicting a single joint distribution over valid temporal intervals, and curate 90K-scale grounded samples and human-verify 3K-scale benchmark annotations.
Yuhui Zeng, Xin-Yu Mao, Xiaokun Liu et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.