Preprint
Aug 2026
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No
FV-Action, the training-free method built on this analysis, is the strongest training-free result on this benchmark; it surpasses every TVG-trained model evaluated zero-shot on TACoS, and improves over direct prediction on ActivityNet Captions and QVHighlights, with no temporal supervision at any stage.
Ji Huang, Barry Devereux, Hui Wang
· 0 citations