Preprint
Aug 2026
COMET: Contrastive Motion-Enhanced Temporal Reasoning for Video Multimodal Large Language Models
COMET is a temporally grounded framework that systematically strengthens video MLLMs through explicit temporal representation, appearance-motion fusion, and direction-aware optimization and achieves consistent overall improvements with a pronounced motion-temporal bias.
Chenghua Zhu, Zhaolu Kang, Qifan Shi et al.
· 0 citations