Preprint
Jul 2026
Mitigating Modality and Language-Style Gaps for Zero-Shot Video Moment Retrieval
This work proposes Self-SiMS, a self-similarity-based Moment Proposal and Scoring that exploits intrinsic relationships within videos, enabling robust span generation and scoring and introduces a query-aware MLLM-based reasoning stage to further sharpen alignment between text and video.
Jihyun Lee, Cheol-Ho Cho, Woojin Jun et al.
· 0 citations