Preprint
Sep 2026
S$^2$Prune: Spatially Structured Visual Token Pruning for Multimodal Large Language Models
S$^2$Prune is proposed, a training-free pruning method that preserves spatial coverage while adapting token density to local image structure and achieves the highest average accuracy among the evaluated training-free pruning methods.
Yuan-Yuan Jia, Shun-Pu Tang, Qian-Qian Yang
· 1 citation
· ⚡1