General-purpose agents increasingly write code, use tools, and complete complex digital tasks, raising the question of how far these capabilities carry into the physical world. To investigate this, we introduce RobotWorld, a challenging simulation testbed for robot use: turning instructions and observations into physic...
Zhiqin Yang, Chen-Xin Li, Xiao-Meng Hu et al.· 0 citations
The CVPR 2026@AdvML Workshop Challenge on adversarial multimodal attacks against autonomous-driving VLAs is presented, providing a practical reference for future robustness evaluation and defense design in multimodal autonomous-driving systems.
Tian-Yuan Zhang, Zonglei Jing, Jiangfan Liu et al.· arXiv.org· 0 citations
Omni-RewardBench is introduced, the first benchmark for comprehensive evaluation of ORMs across modalities and demonstrates that current OLLMs fall short as reward models, revealing several common failure modes such as perception failure, modality dominance failure, and cross-modal fusion failure.
Chi-Min Chan, Yujin Zhou, Pengcheng Wen et al.· Annual Meeting of the Associ...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.