This work shows that paraphrase robustness is a core requirement for reliable VLM-based reward modeling in robotics and introduces ROBORMBENCH, a benchmark with 2,390 real-robot trajectories, ground-truth progress labels, and 21,673 verified paraphrases spanning lexical, syntactic, and action-goal rewrites.
Abstract
Vision-language models are increasingly used as reward functions for robotic learning, but this role requires paraphrase invariance: the same trajectory should receive the same reward under semantically equivalent goal descriptions. We show that current VLM reward models often violate this property. Paraphrasing the instruction alone can substantially change predicted progress scores, and can even flip identical robot behavior between failure and success. To measure this failure mode, we introduce ROBORMBENCH, a benchmark with 2,390 real-robot trajectories, ground-truth progress labels, and 21,673 verified paraphrases spanning lexical, syntactic, and action-goal rewrites. Across proprietary and open-source VLMs, paraphrase-induced instability is widespread and severe, grows under more divergent rewrites, and is not reliably reduced by scale or explicit reasoning. Dedicated reward models trained with trajectory-grounded supervision are substantially more stable. These results show that paraphrase robustness is a core requirement for reliable VLM-based reward modeling in robotics.
This work introduces Dream2Reward, which learns a language-conditioned successful latent transition field from positive demonstrations that provides stronger success-failure separation and more informative feedback on low-quality behavior than progress-based alternatives.
Hao-Yu Zhang, Ze-Cui Zeng, Bin Wang et al.· 3 citations
This work introduces Robo-Dopamine 2.0, a history- and OOD-aware process reward model with a pairwise prediction interface that combines history-conditioned pairwise rewards that use source-aligned reference panels for synthetic OOD queries and observed rollout history for online queries, while preserving the queried e...
Yijie Xu, Hao-Peng Jin, Run Zhou et al.· 2 citations
Despite rapid progress, generalist robot policies remain brittle on complex, long-horizon tasks that comprise multiple stages or require repeated attempts and deliberation on the same underlying stage before success. Q-value functions can improve these policies by ranking candidate actions or guiding policy improvement...
Saksham Singh, Zhe-Yuan Hu, Max Sobol Mark et al.· 0 citations
Across four RoboTwin tasks spanning different horizons and coordination patterns, Prism-GRPO improves success and quality at matched rollout budgets and reaches target success rates with up to 56% fewer rollouts.
Zeyun Deng, Yuzhe Lu, Ya-Wei Wang et al.· 1 citation
General-purpose vision-language models offer a promising way to zero-shot robot control: \gptastra{} excels at open-ended and language- or image-conditioned manipulation but remains substantially weaker on high-precision and long-horizon tasks. We introduce \emph{RoboICL}, an in-context robot-control framework that nar...
Fang-Cheng Liu, Ye-Qing Shen, An-Da Cheng et al.· 0 citations
Robot demonstration datasets used to train vision-language-action policies can contain a subtle but harmful failure mode: trajectories that are behaviorally correct but paired with the wrong language instruction. We study post-hoc auditing of these Instruction–Trajectory Mismatches (ITMs). Unlike failed rollouts, ITMs...
Simon Holk, Ryosuke Takanami, T. Matsushima et al.· IEEE Robotics and Automation...· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduJun 26, 2026
To help robots do chores in places like homes and factories, a new approach from MIT uses one language model to clarify users’ instructions, then another to ignore irrelevant info.