Skip to content

LENS: Adaptive Spatio-Temporal Zooming for Keyframe Sampling in Long-Form Videos

Jul 2026 · arXiv.org · Vol abs/2607.25125 · 2 citations · 62 references
Computer Science

TL;DR

LENS is a training-free keyframe sampling framework that dynamically decides when to zoom in for fine-grained details and when to zoom out for broader context based on the text query, enabling the model to reason across multiple granularities while capturing both high-fidelity details and long-range context.

Abstract

Despite rapid progress in Multi-modal Large Language Models (MLLMs), understanding long-form videos is still bottlenecked by limited context windows. While recent keyframe sampling methods attempt to mitigate this by distilling video inputs into a compact set of query-relevant frames, navigating the vast spatio-temporal search space remains challenging, as spatial detail and temporal coverage often conflict. To address this, we introduce LENS, a training-free keyframe sampling framework that dynamically decides when to zoom in for fine-grained details and when to zoom out for broader context based on the text query. Concretely, LENS adaptively allocates a limited frame budget between spatial zoom-ins, which highlight query-relevant regions within individual frames, and temporal zoom-outs, which expand the temporal scope through multi-frame aggregation, enabling the model to reason across multiple granularities while capturing both high-fidelity details and long-range context. Across diverse long-form video benchmarks, LENS consistently outperforms prior state-of-the-art keyframe sampling methods and delivers substantial gains over uniform sampling, improving Video-MME accuracy from 53.3% to 60.7% with Qwen2.5-VL.Code is available at https://github.com/zhangce01/LENS.

View source

Similar papers

Preprint Aug 2026

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding

A fine-grained dynamic visual selection framework named EviSelect, grounded in the target MLLM internal attention evidence, which efficiently approximate attention maps of the target MLLM using highly compressed visual inputs and sparse attention, well-aligned to the full counterpart.

Bo Zhang, Wenxin Wang, Feng Chen et al. · 0 citations
Preprint Sep 2026

Intrinsic Temporal Adaptation of CLIP for Partially Relevant Video Retrieval

Partially Relevant Video Retrieval (PRVR) aims to retrieve untrimmed videos that contain moments relevant to a text query. Since the target moment occupies only a portion of the video, PRVR requires retrieval based on fine-grained understanding beyond coarse video-level matching. However, existing methods often rely on...

Hyun Seok Seong, Woojin Jun, Subeen Lee et al. · 0 citations
#large language models Preprint Sep 2026

ViTAL-X: Video-Text Alignment with Cross-Modal Temporal Edits

ViTAL-X, a lightweight model that equips frozen image-text backbones with temporal awareness while preserving their foundational spatial knowledge, achieves state-of-the-art performance and demonstrates that targeted, high-quality temporal alignment provides a highly efficient alternative to pure scaling.

V. SethuramanT., Savya Khosla, O. Susladkar et al. · 0 citations
Preprint Aug 2026

Temporal Tree of Thought: Reasoning-Guided Visual Cue Search for Long-Video Understanding

The proposed Temporal Tree of Thought T^3, a training-free framework for adaptive coarse-to-fine long-video understanding, constructs a question-agnostic hierarchical temporal tree via recursive temporally constrained clustering, where each node represents a contiguous segment with an informative key frame.

Ziling Huang, Shin'ichi Satoh · 0 citations
Preprint Aug 2026

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs

VideoRouter (VR) is proposed that rethinks long-video understanding as coordinating complementary evidence views rather than selecting a single subset of frames, and introduces a verification-guided router to determine which view is better supported by the selected evidence and select the final answer.

Ziling Huang, Yuki M. Asano, Shin'ichi Satoh · 0 citations
#artificial intelligence Preprint Sep 2026

WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory

Video world models enable interactive exploration of dynamic environments, yet struggle to respect prior observations over long horizons and across viewpoints. We present WorldCrafter, a video world model that learns a camera-queryable implicit 3D-aware memory for this purpose. The key insight is to let the requested v...

Wang-Bo Yu, Kunhao Liu, Wen-Bo Hu et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.