This study introduces ClustRS, a two-part, training-free algorithm for robust token pruning, demonstrating a simple yet powerful alternative to both score-only and diversity-only pruning rules, paving the way for compute-efficient and noise-resilient VLM deployment.
Abstract
Recent Visual-Language Models (VLMs) have enhanced the capabilities of pre-trained LLMs by adding vision tokens alongside text, with approaches like LLaVA showing impressive results. However, the computational burden of processing up to 576 or 729 visual tokens makes edge deployment challenging. While various token pruning techniques require retraining, some are training-free and thus can easily adapt to architecture changes. We introduce ClustRS, a two-part, training-free algorithm for robust token pruning. Its first component is an attention-weighted, clustering algorithm that selects representative tokens from each semantic cluster. The second component, Residual Shrinkage, is a one-pass denoising step on the selected tokens. These training-free lightweight steps make LLaVA ready for real-world data, improving robustness to a wide range of image-noise types and intensities. Experimental results on the ScienceQA-IMG and MM-VET benchmarks show our method outperforms attention- and diversity-based methods by up to 20\% under extreme noise and token conditions (reducing tokens by 97\%, down to 16 tokens) on LLaVA 1.5 7b and achieves exceptional results on LLaVA-OneVision, where we match baseline performance with fewer than one-third of their tokens under mild noise conditions. Our study demonstrates a simple yet powerful alternative to both score-only and diversity-only pruning rules, paving the way for compute-efficient and noise-resilient VLM deployment.
A training-free, Adaptive Visual Token Pruning (AVTP) framework, applicable to diverse LVLM architectures is proposed, and adaptive pruning ratios in multi-image contexts where images of higher importance retain proportionally more tokens are implemented.
Rongyang Zhang, Cheng-Qiang Lu, Cong Li et al.· 0 citations
Results on the UC Merced (UCM) and NWPU benchmarks indicate that SE-CLIP significantly outperforms existing semi-supervised approaches and provides a viable solution for adapting VLMs to the remote sensing domain with minimal human intervention.
M. L. Mekhalfi, M. M. Al Rahhal, Y. Bazi et al.· IEEE Geoscience and Remote S...· 0 citations
SinkPruner is proposed, a training-free visual token pruning framework for efficient MLLM inference that follows a coarse-to-fine design with two key modules: a visual sanitizer that filters high-norm redundancies and alleviates attention sink and attention dispersion, and a text-guided pruner that further retains toke...
Shi-Yu Li, Zi-Yuan Hu, Shijia Huang et al.· 1 citation
Vision-language pre-training has reshaped image clustering, giving rise to language-assisted image clustering (LaIC), which leverages textual semantics to complement visual representations. Despite the rapid proliferation of LaIC methods, it remains unclear how much LaIC has actually advanced image clustering, as exist...
This work proposes a two-stage adaptive token pruning strategy specifically designed for video processing that improves accuracy by +7\% on a video captioning benchmark at 10% token retention, while reducing computation TFLOPs by 95\%.
Paribesh Regmi, Qingshuang Chen, Chi Zhang et al.· 1 citation
ClusterAttention is the first training-free method that can be successfully applied in the setting of unstructured input and a single forward pass, and is the first training-free method that can be successfully applied in the setting of unstructured input and a single forward pass without offline calibration.
Kasper Nordenram, Amelie Dittmann· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.