Super-Image Reranking: Lightweight Temporal Context Aggregation for Ad-hoc Video Search
Ad-hoc video search retrieves relevant video shots from large-scale collections using free-form textual queries. Pretrained vision-language models provide strong frame-level matching, but max-similarity aggregation can miss semantic evidence distributed across neighboring frames. We propose a lightweight retrieve-then-...