S$^2$Prune is proposed, a training-free pruning method that preserves spatial coverage while adapting token density to local image structure and achieves the highest average accuracy among the evaluated training-free pruning methods.
Abstract
Visual token pruning reduces the inference overhead of multimodal large language models (MLLMs) by retaining only a subset of visual tokens. Existing methods usually select tokens based on importance or redundancy. However, we observe that these criteria produce stable spatial biases across inputs and do not always outperform simple Uniform Grid sampling, highlighting the value of broad spatial coverage. Motivated by this, we propose S$^2$Prune, a training-free pruning method that preserves spatial coverage while adapting token density to local image structure. We first divide the image into regions and assign at least one token to each region to preserve coverage. The remaining token budget is then distributed according to Laplacian variation, giving more tokens to regions with richer structure. We then use Early Representation Change (ERC), computed from the first decoder block, to select representative tokens within each region. We evaluate S$^2$Prune across diverse settings and two MLLM architectures. On Qwen2.5-VL-7B-Instruct, it achieves the highest average accuracy among the evaluated training-free pruning methods. With only 32 of the original 576 visual tokens, it still retains 79.3% of the full-model performance. Code is available at https://github.com/yuanyuanjia71-spec/S2Prune.
Visual prefixes account for a major portion of the per-layer computation in multimodal large language models (MLLMs), making visual-token pruning a direct approach to accelerating inference. Existing top-K methods typically evaluate tokens independently and apply a uniform budget to all inputs, overlooking both selecti...
Han-Sen Zhang, Lan He, Min Yao et al.· 0 citations
A training-free, Adaptive Visual Token Pruning (AVTP) framework, applicable to diverse LVLM architectures is proposed, and adaptive pruning ratios in multi-image contexts where images of higher importance retain proportionally more tokens are implemented.
Rongyang Zhang, Cheng-Qiang Lu, Cong Li et al.· 0 citations
SinkPruner is proposed, a training-free visual token pruning framework for efficient MLLM inference that follows a coarse-to-fine design with two key modules: a visual sanitizer that filters high-norm redundancies and alleviates attention sink and attention dispersion, and a text-guided pruner that further retains toke...
Shi-Yu Li, Zi-Yuan Hu, Shijia Huang et al.· 1 citation
CoverPruner is proposed, a training-free pruner that asks the complementary demand-side question: after a token is removed, which surviving original token represents it for the target VLM?
Qin Zhu, Wei-Hang You, Han-Qi Jiang et al.· 2 citations
This work proposes $\delta$-Vision, which replaces repeated Transformer evolution of visual tokens with lightweight low-rank adapters that construct layer-wise visual memories while preserving all visual tokens for text retrieval, and achieves higher accuracy than visual token pruning baselines at comparable or lower c...
Jing-Di Lei, Junxian Li, Di Zhang et al.· 0 citations
Existing VLA pruning strategies primarily select individual visual tokens according to task-level semantic relevance, while overlooking the spatial information required for robotic manipulation. To examine this limitation, we construct a simple Stride baseline that uniformly samples tokens along the flattened one-dimen...
Jia-Yu Chen, Shuyong Gao, Jing-Kai Jia et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.