Preprint
Jul 2026
SepPrune:A Separator-based Pruning Framework for Efficient Multimodal Large Language Models
It is observed that attention scores from both vision and text tokens peak at modality separator tokens, suggesting that these separators bridge the two modalities and proposes SepPrune, an efficient, training-free, plug-and-play pruning method that uses the separator token as a unified query to rank and select informative vision tokens.
Yucheng Wang, Qihui Zhu, Yang Liu et al.
· 0 citations