Recent text-to-image models have made substantial progress in photorealism, aesthetics, and text-image alignment. Yet visually appealing images can still violate real-world plausibility, exhibiting malformed object structures, impossible anatomy, physically implausible interactions, or inconsistent spatial relationship...
NavMCP is introduced, an agentic scaffolding framework that couples a VLM reasoning agent with an NFM executor for long-horizon exploration that achieves state-of-the-art results on HM-EQA, MT-HM3D, and EXPRESS-Bench.
Zi-Xing Lei, Geng-Ze Zhou, Xiong-Hui Chen et al.· 0 citations
The results show that a general-purpose model can already achieve competitive embodied control without a navigation policy, and term this organization agentic embodied control: the reasoning model directly steers every action, keeping reasoning and control aligned.
Referring Video Object Segmentation (RVOS) aims to segment the target objects specified in human instructions. Previous approaches typically rely on explicit human instructions that contain target categories or salient appearance descriptions. These approaches tend to fail when the instructions require temporal video u...
Yanyan Shao, Shuting He, Gengze Zhou et al.· IEEE Transactions on Image P...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.