Results support the effectiveness of caption memory for episodic reasoning over long egocentric video in wearable assistants with bounded frame budgets, growing visual-token costs, and long-context retrieval failures.
Dingli Liang, Yi Xie, Yu-Kai Huang et al.· 0 citations
This work proposes Caption-once, Frames-onDemand (CFD), a budget-aware edge-cloud agentic framework that turns visual access into a first-class, query-conditioned cost, capping per-query frame consumption regardless of video length.
Wei-Tong Cai, Hang Zhang, Yu-Kai Huang et al.· 2 citations
AVE-Agent is proposed, a modular agent framework that decomposes complex instructions into dependent subtasks and iteratively improves editing results through self-reflection and evaluator feedback, and improves instruction execution, Fidelity Preserving, and audio-visual alignment in joint editing while maintaining co...