Skip to content
Preprint

Shallow to Deep: Aligning Token Pruning with Stage-wise Roles in LVLMs

Sep 2026 · 0 citations · 33 references
Computer Science

TL;DR

STD is proposed, a hierarchical token pruning framework that adapts token selection mechanisms to the functional role of each network stage, and introduces a Stability-Adaptive Trigger in deep layers to execute pruning only during semantically stable phases.

Abstract

Large Vision-Language Models (LVLMs) incur high computational costs from redundant visual tokens. Although training-free attention-based multi-layer pruning in the vision encoder stage has been explored as an effective strategy, we find that pruning in shallow layers consistently degrades performance. In this paper, we aim to understand this problem and seek a solution. By analyzing attention patterns across network depth, we find that shallow layers primarily function as edge detectors with chaotic attention maps, while deeper layers transition through local subject recognition and unstable semantic aggregation. To address the misalignment between pruning strategies and network stages, we propose STD, a hierarchical token pruning framework that adapts token selection mechanisms to the functional role of each network stage. STD employs High-Frequency Spectral Analysis in shallow layers to deterministically preserve structural edges, uses Gaussian-Smoothed Attention in intermediate layers to maintain spatial coherence, and introduces a Stability-Adaptive Trigger in deep layers to execute pruning only during semantically stable phases. Extensive experiments show that STD outperforms state-of-the-art pruning methods by 1.1% on LLaVA-1.5-7B with 88.9% token reduction, while also being plug-and-play and highly effective when combined with other methods, and by 2.1% on LLaVA-NeXT-7B with 94.4% reduction, delivering a 3.9x speed-up in the prefilling stage. Our code will be released at https://github.com/Twilight03/STD.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

AdaVSkip: Adaptive Visual Token Skipping Across Layers For Efficient MLLMs Inference

AdaVSkip is proposed, which equips each layer with two lightweight routers that independently determine whether visual tokens pass through by or skip the self-attention and MLP modules, and maintains strong task performance with substantially less computation.

Yu-Yao Sun, Tao Deng, Shuang-Hua Li et al. · 0 citations
#artificial intelligence Open access Sep 2026

SPIDER: Multi-Layer Semantic Token Pruning and Adaptive Sub-Layer Skipping in Multimodal Large Language Models.

Multimodal Large Language Models face significant efficiency challenges that stem from two distinct yet coupled sources: data redundancy and computational redundancy. While most methods focus on data redundancy by pruning visual tokens from the output of the visual encoder or computing redundancy in LLM decoders using...

Tian-Xiang Chen, Zhentao Tan, Zi Ye et al. · 0 citations
Preprint Sep 2026

Layer-Aware Position Embeddings for Visual Token Pruning in Multimodal Large Language Models

Multimodal large language models (MLLMs) incur substantial computational overhead due to the reliance on hundreds of visual tokens to represent images. While token pruning has emerged as a promising approach to reduce the inference cost of MLLMs, existing methods typically reassign position embeddings to the retained t...

Ya-Hong Wang, Zhang-Kai Ni, Jun-Cheng Wu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

STAR-Pro: Stage-Wise Token Adaptive Reduction with Progressive Refinement for Efficient Large Vision-Language Models

Large vision-language models (LVLMs) achieve strong multimodal understanding, but the hundreds to thousands of visual tokens they process impose substantial computational overhead, motivating training-free visual token pruning. In this work, we conduct two complementary analyses of visual token pruning. First, we measu...

Yi-Chen Guo, Ting-Hao Wang, Qi-Zhe Zhang et al. · 1 citation
Preprint Aug 2026

E2S-Pruner: Progressive Two-Stage Evidence Fusion for Visual Token Pruning in Vision-Language Models

A spatial novelty constraint is introduced that promotes coverage of distinct image regions and prevents the retained tokens from concentrating in a few locally salient areas and prevents the retained tokens from concentrating in a few locally salient areas in E2S-Pruner.

Taoyu Qian, Qi Wang, D. Shi et al. · 1 citation
#artificial intelligence Preprint Sep 2026

VPRune: Efficient Training-free Pre-LLM Visual Token Pruning

Visual token pruning is a promising approach to reducing the inference cost of large vision-language models (LVLMs), yet aggressive token reduction often causes substantial performance degradation. We identify three key factors behind this degradation: text-guided selection bias, information loss from discarded tokens,...

Guang-Chuan Lv, Dian-Xing Shi, Ding-Jie Fu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.