Preprint
Jul 2026
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
OSReward is introduced, a realistic, high-quality benchmark that evaluates VLM judges on CUA trajectories, and an open corpus of reasoning-annotated trajectory judgments for the CUA community, to close the gap in reliable CUA reward at scale.
Qiushi Sun, Kanzhi Cheng, Yian Wang et al.
· 1 citation
· ⚡1