Omni-RewardBench is introduced, the first benchmark for comprehensive evaluation of ORMs across modalities and demonstrates that current OLLMs fall short as reward models, revealing several common failure modes such as perception failure, modality dominance failure, and cross-modal fusion failure.
Chi-Min Chan, Yujin Zhou, Pengcheng Wen et al.· Annual Meeting of the Associ...· 0 citations
FISA is proposed, a framework for MLLM self-improvement that constructs augmented images from the model's own failure cases that generates visually challenging yet answer-preserving image complications, verifies their utility through self-examination, and applies dual fidelity filtering to avoid semantic distortion.
Chunyang Jiang, Pingping Zhang, Yuzhi Zhao et al.· 0 citations