Omni-RewardBench is introduced, the first benchmark for comprehensive evaluation of ORMs across modalities and demonstrates that current OLLMs fall short as reward models, revealing several common failure modes such as perception failure, modality dominance failure, and cross-modal fusion failure.
Chi-Min Chan, Yujin Zhou, Pengcheng Wen et al.· Annual Meeting of the Associ...· 0 citations
AgentGym2 is presented, a new evaluation framework with task instances grounded in real-world end-to-end working demands that measures agents'ability to execute end-to-end procedures, discover tools via exploration, compose tools for unseen tasks, and remain robust to noisy and underspecified information.
Zhiheng Xi, Dingwen Yang, Jiaqi Liu et al.· Annual Meeting of the Associ...· 1 citation