Multimodal agents commonly generate free-form reasoning before each action. For small models, limited model capacity can result in lengthy reasoning that provides little useful guidance for action generation while incurring substantial inference cost. To address this challenge, we introduce Selection-based Structured R...
Feiyu Zhu, Xiao-Yu Zhu, Ji-Qian Yang et al.· 0 citations
Structured policies improve efficiency, robustness, and interpretability in imitation learning by introducing task-specific inductive bias, but existing structure generation methods rely either on extensive human input or on static domain knowledge encoded in LLMs, which may be inconsistent with the expert demonstratio...
Feiyu Zhu, Qi Xu, Zhi-Fei Deng et al.· 0 citations
This work shows that paraphrase robustness is a core requirement for reliable VLM-based reward modeling in robotics and introduces ROBORMBENCH, a benchmark with 2,390 real-robot trajectories, ground-truth progress labels, and 21,673 verified paraphrases spanning lexical, syntactic, and action-goal rewrites.
Wonje Jeung, Sangyeon Yoon, Hyesoo Hong et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.