Skip to content

Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward Models

Sep 2026 · 0 citations · 57 references
Computer Science

TL;DR

This work shows that paraphrase robustness is a core requirement for reliable VLM-based reward modeling in robotics and introduces ROBORMBENCH, a benchmark with 2,390 real-robot trajectories, ground-truth progress labels, and 21,673 verified paraphrases spanning lexical, syntactic, and action-goal rewrites.

Abstract

Vision-language models are increasingly used as reward functions for robotic learning, but this role requires paraphrase invariance: the same trajectory should receive the same reward under semantically equivalent goal descriptions. We show that current VLM reward models often violate this property. Paraphrasing the instruction alone can substantially change predicted progress scores, and can even flip identical robot behavior between failure and success. To measure this failure mode, we introduce ROBORMBENCH, a benchmark with 2,390 real-robot trajectories, ground-truth progress labels, and 21,673 verified paraphrases spanning lexical, syntactic, and action-goal rewrites. Across proprietary and open-source VLMs, paraphrase-induced instability is widespread and severe, grows under more divergent rewrites, and is not reliably reduced by scale or explicit reasoning. Dedicated reward models trained with trajectory-grounded supervision are substantially more stable. These results show that paraphrase robustness is a core requirement for reliable VLM-based reward modeling in robotics.

View source

Similar papers

Preprint Aug 2026

Dream2Reward: Transition-Alignment Reward Models from Positive Demonstrations for Robotic Manipulation

This work introduces Dream2Reward, which learns a language-conditioned successful latent transition field from positive demonstrations that provides stronger success-failure separation and more informative feedback on low-quality behavior than progress-based alternatives.

Hao-Yu Zhang, Ze-Cui Zeng, Bin Wang et al. · 3 citations
Preprint Aug 2026

Robo-Dopamine 2.0: History-Conditioned and OOD-Aware Process Reward Modeling for Robotic Manipulation

This work introduces Robo-Dopamine 2.0, a history- and OOD-aware process reward model with a pairwise prediction interface that combines history-conditioned pairwise rewards that use source-aligned reference panels for synthetic OOD queries and observed rollout history for online queries, while preserving the queried e...

Yijie Xu, Hao-Peng Jin, Run Zhou et al. · 2 citations
Preprint Sep 2026

SeeQ: Training Generalist Value Functions for Long-Horizon Robotic Manipulation

Despite rapid progress, generalist robot policies remain brittle on complex, long-horizon tasks that comprise multiple stages or require repeated attempts and deliberation on the same underlying stage before success. Q-value functions can improve these policies by ranking candidate actions or guiding policy improvement...

Saksham Singh, Zhe-Yuan Hu, Max Sobol Mark et al. · 0 citations
#machine learning Preprint Sep 2026

RoboICL: Embodied In-Context Learning with GPT-6 Astra

General-purpose vision-language models offer a promising way to zero-shot robot control: \gptastra{} excels at open-ended and language- or image-conditioned manipulation but remains substantially weaker on high-precision and long-horizon tasks. We introduce \emph{RoboICL}, an in-context robot-control framework that nar...

Fang-Cheng Liu, Ye-Qing Shen, An-Da Cheng et al. · 0 citations
Open access Aug 2026

Auditing Instruction–Trajectory Mismatches in Multimodal Robot Demonstrations

Robot demonstration datasets used to train vision-language-action policies can contain a subtle but harmful failure mode: trajectories that are behaviorally correct but paired with the wrong language instruction. We study post-hoc auditing of these Instruction–Trajectory Mismatches (ITMs). Unlike failed rollouts, ITMs...

Simon Holk, Ryosuke Takanami, T. Matsushima et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.