Preprint
Jul 2026
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models
An attention-free and lightweight token reduction framework as a plug-and-play module for VLMs, which preserves both important and diverse tokens to produce a compact visual representation, and achieves a favorable accuracy-efficiency trade-off.
Xuanyi Hao, Zuoyuan Zhang, Zhibo Wang et al.
· 0 citations