Super-Image Reranking: Lightweight Temporal Context Aggregation for Ad-hoc Video Search
Abstract
Ad-hoc video search retrieves relevant video shots from large-scale collections using free-form textual queries. Pretrained vision-language models provide strong frame-level matching, but max-similarity aggregation can miss semantic evidence distributed across neighboring frames. We propose a lightweight retrieve-then-rerank framework that augments a frame-level image-text baseline with super-image reranking. For top-ranked candidate shots, consecutive frames are arranged into multi-frame grid images and matched with the query using the same backbone. The super-image score is fused with frame-level refinement and merged with the original retrieval score through query-wise normalization. Experiments on TRECVID AVS show that the proposed method improves infAP from 0.2160 to 0.2414 over a BEiT-3 frame-level max-similarity baseline. Ablations confirm that super-image context and frame refinement provide complementary signals. Zero-shot experiments on MSR-VTT with CLIP backbones further show consistent gains over frame-max retrieval, suggesting that super-image reranking is a simple and effective strategy for improving text-to-video retrieval without dedicated video-text training.