From Saliency to Discriminability: Rank-Preserving Visual Token Pruning for VLM Rerankers
Large vision-language models used as listwise rerankers must jointly process visual tokens from tens of candidates per query, making token pruning essential for practical deployment. Existing pruning methods retain tokens by attention saliency, yet we show that saliency is systematically misaligned with ranking contrib...