Recent advances in video generative models have enabled high-fidelity, temporally coherent video generation. However, these models often struggle to satisfy prompts requiring specialized knowledge, specific identities, physical consistency, or ordered events. In this paper, we present VideoGen-Agent, a multimodal agent...
Bin-Xu Li, Hao-Yi Duan, Yu-Hui Zhang et al.· 0 citations
Referring remote sensing image segmentation (RRSIS) aims to localize and segment specific geospatial targets in images guided by natural language expressions. Although existing methods have achieved promising progress in vision-language alignment, they still suffer from three major limitations in complex scenarios, inc...
Chun-Yuan Li, Yu-Xiang Xie, Qian-Qi Lu et al.· 2026 12th International Conf...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.