Long video understanding often behaves like a visual needle-in-a-haystack problem: query-relevant evidence is sparsely distributed across long temporal spans, while packing dense frames into a single VLM context incurs \textit{context rot} and high cost. Existing video agents often rely on query-agnostic offline prepro...
Sen Yang, Bo-Qiang Duan, Jing Yang et al.· 0 citations
Multimodal Large Language Models (MLLMs) have been growing the capability for scientific writing and collaboration. For example, OpenAI Prism is a free workspace for scientific writing and collaboration. One important feature in Prism is turning scientific diagrams directly into LaTeX TikZ code. In this paper, we build...
Wei-Hao Bo, Shan Zhang, Yanpeng Sun et al.· 1 citation· ⚡1
Video object removal is a fundamental yet challenging task in video editing. Despite recent progress, existing methods typically fall into two categories. Traditional approaches based on optical flow or attention mechanisms often introduce noticeable artifacts and yield unnatural results. In contrast, diffusion-based m...
Zi-Zhao Chen, Ping Wei, Guang Dai et al.· 1 citation
GroundShot is presented, a training-free, model-agnostic agentic framework for entity-grounded multi-shot generation that improves multi-shot consistency over existing methods while requiring no additional training or model modification.
Yixuan Lai, Tianjia Shao, Kun Zhou et al.· arXiv.org· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.