Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-reference image-grounded video captioning, a new task requiring factual video descriptions with phrase-level reference grounding, and propos...
Tengfei Liu, Yang Shi, Yuran Wang et al.· 0 citations
The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (MLLMs) on complex tasks, rather than merely equipping them with a sophisticated yet inefficient reasoning paradigm. In this work, we rethink agentic visual reasoning through two key dimensions of tool us...
Qixun Wang, Yang Shi, Le-Tian Cheng et al.· arXiv.org· 0 citations
FocusMem is introduced, which separates episodic memory and working memory within a compact latent-memory interface and consistently outperforms a fully matched action-only fixed-memory baseline and prior latent memory adaptations.
Zhuoran Zhang, Bowen Li, Jingcheng Ju et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.