World Action Models (WAMs) have emerged as a promising paradigm for generalizable robot control. Despite the growing number of WAM systems, existing works often introduce unified systems that bundle together multiple design choices, such as architecture and training strategy, making it difficult to isolate individual c...
Chao Tang, Haoqing Wang, Zi-Lang Cen et al.· 0 citations
General-purpose vision-language models offer a promising way to zero-shot robot control: \gptastra{} excels at open-ended and language- or image-conditioned manipulation but remains substantially weaker on high-precision and long-horizon tasks. We introduce \emph{RoboICL}, an in-context robot-control framework that nar...
Fang-Cheng Liu, Ye-Qing Shen, An-Da Cheng et al.· 0 citations
FaithEyes, a multi-agent self-judging framework that uses a VLM to judge whether each process image helps answer the question and designs a multi-agent framework where the model itself serves as a subagent to judge the tool calls from the main agent, eliminating any dependence on external models at inference.
Vision-language models (VLMs) have made rapid progress in visual perception and increasingly support real-world tasks that depend on images. Many such tasks, however, require more than rec- ognizing what an image contains: a model must use visual evidence to make a complete decision whose parts jointly satisfy global c...
Ning-Xin Pan, Han-Yu Li, Ye-Hui Tang· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.