Preprint
Aug 2026
ID-VTG: Image-Disambiguated Video Temporal Grounding
The Visually-Guided Disambiguation Aggregation Aggregation (VGD-Agg) framework is proposed, a framework based on a dual-branch fast-slow architecture that enhances discriminability via two learnable tokens and achieves state-of-the-art results on the proposed benchmarks.
Minghang Zheng, Jing Wei, Hong-Yi Yang et al.
· 0 citations