A wearable assistant should both answer questions about its visual history and recognize when that history is useful to the present situation. Existing video-memory systems primarily support question-conditioned recall, whereas proactive assistants typically use separate memory and control mechanisms. We introduce GROVE, a training-free framework that supports both behaviors with one memory grown causally from a continuous video stream. GROVE retains fine-grained perceptual evidence and incrementally consolidates it into time-stamped moments, coherent episodes, and recurring cross-day patterns. Each stratum is paired with a scale-native retrieval skill for locating an observation, replaying an activity, or traversing long-range regularities. Reactive QA and proactive assistance share this memory and access interface, differing in whether retrieval is initiated by a user query or the current situation. Across multiple benchmarks including the challenging MM-lifelong and EgoServe, GROVE achieves the best results among the compared methods. Controlled ablations show that the temporal strata and their access skills are complementary, with patterns providing the largest benefit when evidence spans multiple days. Code will be available at https://github.com/SitongGong/GROVE.
Sitong Gong, Caixin Kang, Tianyu Yan et al.· 0 citations
This work demonstrates that the aligned LLM with a general-purpose vision encoder can effectively enhance downstream VQA performance with task-specific encoders, and investigates several alignment strategies between the aligned LLM and new task-specific encoders.
Jiazuo Yu, Yunzhi Zhuge, Lu Zhang et al.· International Journal of Com...· 0 citations
ERA establishes logit-preserving visual token pruning as a principled framework for efficient MLLMs, unifying theoretical foundation, algorithmic design, and practical deployment.
Yuhao Wang, Mu Qiao, Haiwen Diao et al.· arXiv.org· 0 citations