Skip to content

Caption-once, Frames-on-Demand: Visual-Need Routing for Budget-Aware Agentic Long Video Understanding

Sep 2026 · 2 citations · 49 references
Computer Science

TL;DR

This work proposes Caption-once, Frames-onDemand (CFD), a budget-aware edge-cloud agentic framework that turns visual access into a first-class, query-conditioned cost, capping per-query frame consumption regardless of video length.

Abstract

Long-video understanding on edge devices must reason over hours of content under tight compute and bandwidth budgets. Subsampling visual tokens loses temporal structure, while text-only video memories lose fine-grained visual attributes. We observe a visual-textual duality: language memories carry long-range temporal structure better than dense frames, while pixels remain decisive for attribute-level perception. Building on this insight, we propose Caption-once, Frames-onDemand (CFD), a budget-aware edge-cloud agentic framework. The edge runs a single offline captioning pass that builds a dual-track narrative index, an event-level story skeleton plus a clip-level micro-log, cached and reused across queries without re-captioning. At query time, a cloud-side MLLM reasons over the index in a story-first loop centered on a lightweight Visual-Need Router: a per-query gating module that triggers bounded keyframe retrieval only for perceptual questions (appearance, on-screen text, attribute disambiguation) and keeps temporal-structural questions in language space. The router turns visual access into a first-class, query-conditioned cost, capping per-query frame consumption regardless of video length. Experiments on long-video benchmarks demonstrate strong accuracy-efficiency trade-offs while substantially reducing online visual processing.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Online Video Agent Harness for Long Video Understanding

Long video understanding often behaves like a visual needle-in-a-haystack problem: query-relevant evidence is sparsely distributed across long temporal spans, while packing dense frames into a single VLM context incurs \textit{context rot} and high cost. Existing video agents often rely on query-agnostic offline prepro...

Sen Yang, Bo-Qiang Duan, Jing Yang et al. · 1 citation
#computer vision Review Sep 2026

Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs

This survey covers inference-efficiency mechanisms for visual and audiovisual VideoLLMs that report concrete reductions in parameter count, FLOPs per input, latency, memory, or visual and audio token count, and organize methods by the pipeline stage at which they act.

Killian Steunou, Yannis Tevissen, M. El Yacoubi · 0 citations
#artificial intelligence Preprint Sep 2026

KeyRec: Bounded Visual Memory for Streaming and Long-Video Understanding

Vision-language models are increasingly used to understand long videos and continuous streams. However, dense visual tokens accumulate with video duration, making long-context inference prohibitively expensive. Existing training-free visual-token selection methods reduce this cost by retaining informative tokens, but m...

Zi-Han Chen, Xuejian Rong, Xiao-Juan Wang et al. · 0 citations
Preprint Sep 2026

CoFiE: Coarse-to-Fine Evidence Selection for Efficient Streaming Video Understanding

Streaming video understanding requires Vision Language Models (VLLMs) to process growing video streams and answer user questions under tight latency constraints. Existing methods improve efficiency through token pruning and memory-bank schemes, but mainly reduce visual tokens after visual encoding. Consequently, downst...

Jing-Chi Jiang, Yi-Ran Ling, Ruo-Nan Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

PREM: Prefix-Steered Recurrent Memory for Long-Video Understanding

Prefix-Steered Recurrent Memory (PREM), a memory-token-free framework for frozen vision-language models (VLMs) that separates video ingestion from query answering, consistently outperforms frozen baselines at every evaluated visual budget.

Si-Ru Zhong, Qiong-Yan Wang, Xiao-Hui Lv et al. · 0 citations
Preprint Aug 2026

Dynamic Hub-and-Spoke Memory for Streaming Video Understanding

Dynamic Hub-and-Spoke Memory is proposed, a training-free framework that represents distant history as structured textual memory while preserving the recent frames as visual tokens for fine-grained perception in streaming video understanding.

Xin-Ru Jiang, Lin Zhao, Xi Xiao et al. · 4 citations

Related blog posts

Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.