Existing long-video agents acquire evidence through one uniform behavior, ignoring whether the required evidence is concentrated, requires broad occurrence coverage, or must discriminate competing hypotheses---which can cause failure before substantive reasoning begins. Prescribing a fine-grained solution procedure for...
Can-Can Zhang, Bao-Feng Zhang, Xiao-Tian Han et al.· 0 citations
Video language models process videos as dense visual-token sequences with substantial representational redundancy. Compressing these sequences is therefore essential for reducing the visual-token burden on language-model decoding. The central challenge is to preserve visual information dispersed across frames under suc...
Wen-Ti Yin, Xiao-Tian Han, Jun-Yuan Shang et al.· 0 citations
Long-context LLM inference is bottlenecked by quadratic attention computation and growing KV-cache costs. Existing sparse attention and KV-compression methods typically decide which tokens or heads to preserve from runtime attention scores, observation windows, calibration prompts, or learned gates, making head diagnos...
YeHan Yang, Jun-Yuan Shang, Yang Li et al.· 0 citations
This work proposes Test-Time Curriculum (TTC), a simple and model-agnostic framework that adapts a detector on unlabeled test data through curriculum-based self-training and substantially improves overall detection performance under diverse unseen-generator shifts, establishing a practical and effective test-time adapt...
Yiqian Zhang, Zheyuan Gu, Xiangzhao Hao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.