Skip to content

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation

Jul 2026 · arXiv.org · Vol abs/2607.23265 · 1 citation · 36 references
Computer Science

TL;DR

WaveZip is proposed, a joint signal-frequency-domain framework for efficient video inference that requires no task-specific training and can be seamlessly integrated into off-the-shelf LVLMs to boost inference efficiency.

Abstract

Existing Large Vision-Language Models (LVLMs) struggle with long-form video understanding due to the quadratic computational cost of visual tokens. While recent efficient methods attempt to compress tokens via hard pruning or uniform merging, they operate strictly in the spatial feature domain, where robust structural context and discriminative semantic details are inherently entangled. In this work, we propose WaveZip, a joint signal-frequency-domain framework for efficient video inference. Driven by the insight that temporal redundancy resides in low-pass approximation scales while spatial saliency strongly correlates with high-frequency components, WaveZip leverages Discrete Wavelet Transforms (DWT) to disentangle these signals. Temporally, it employs 1D DWT to analyze query-frame relevance, and the resulting high-frequency coefficients are further gated by inter-frame differences, with both signals jointly driving the dynamic allocation of a precise frame-level token budget. Spatially, a 2D DWT decomposes features into low-frequency approximations and high-frequency detail components, where the high-frequency coefficients are modulated within query-salient regions to regulate spatial reconstruction. Importantly, WaveZip requires no task-specific training and can be seamlessly integrated into off-the-shelf LVLMs to boost inference efficiency. Extensive experiments on long video understanding benchmarks demonstrate that WaveZip retains 99.6% of the full performance under an extreme 10x compression ratio, consistently outperforming state-of-the-art methods.

View source

Similar papers

Preprint Aug 2026

Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation

KATok (Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation), a transformer-based VAE that incorporates an adaptive token selector which is jointly learned with latent tokens that achieves strong reconstruction and generation quality at a state-of-the-art compression ratio.

Yeonkyeong Lee, Hyun-Young Go, Jongmin Kim et al. · 0 citations
Sep 2026

AdaCompVL: Adaptive Compression of Spatiotemporal and Cross-Modal Redundancy for Efficient Video-Language Learning.

Multimodal large language models (MLLMs) have recently extended from static image understanding to video comprehension, but representing videos as frame-level token sequences incurs substantial computational overhead. Existing visual token compression methods typically rely on uniform sampling or single-dimension redun...

Jian-Xin Ma, Shi-Bo Jin, Lu-Juan Dang et al. · 0 citations
Preprint Aug 2026

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models

Video large language models (VideoLLMs) achieve strong video understanding performance, but their inference remains expensive due to the large number of redundant spatio-temporal visual tokens in long videos. Existing token pruning methods alleviate this cost by reducing redundant tokens, yet most of them rely on segme...

Meng-Jie Zhang, Qi-Hui Zhu, Tao Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

VideoMM: Adaptive Macro-Micro Inference for Efficient Video MLLMs

Scaling Multimodal Large Language Models (MLLMs) to long-form video understanding is bottlenecked by the explosion of visual tokens, which saturates context windows and incurs prohibitive costs. Current solutions predominantly rely on auxiliary models for token reduction but face a fundamental dilemma: lightweight enco...

Hao-Yu Guo, Yuan Feng, Junlin Lv et al. · 0 citations
Preprint Aug 2026

Aggregating Visual Information with Optimal Transport for VideoLM Token Compression

Video language models process videos as dense visual-token sequences with substantial representational redundancy. Compressing these sequences is therefore essential for reducing the visual-token burden on language-model decoding. The central challenge is to preserve visual information dispersed across frames under suc...

Wen-Ti Yin, Xiao-Tian Han, Jun-Yuan Shang et al. · 0 citations
Preprint Aug 2026

CoANeRV: Coordinate-Aware Token-Space Neural Video Representation

Experiments on diverse video datasets show that CoANeRV consistently improves reconstruction quality over prior feed-forward NeRV and INR baselines, reduces peak memory compared with attention-based coordinate decoders, and provides efficient amortized encoding without per-video optimization.

Jialong Guo, Ke Liu, Meng-Xuan Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.