Preprint
Jul 2026
ReToken: One Token to Improve Vision-Language Models for Visual Retrieval
ReToken is a single learnable embedding trained as an explicit retrieval target that selects a sparse set of query-relevant visual tokens from a pre-filled visual KV cache that yields consistent gains across image and video benchmarks.
Yao Xiao, Reuben Tan, Zhen Zhu et al.
· 0 citations