Preprint
Aug 2026
Multi-Image Visual Token Pruning in Large Visual Language Models
A training-free, Adaptive Visual Token Pruning (AVTP) framework, applicable to diverse LVLM architectures is proposed, and adaptive pruning ratios in multi-image contexts where images of higher importance retain proportionally more tokens are implemented.
Rongyang Zhang, Chengqiang Lu, Cong Li et al.
· 0 citations