Evaluating world models requires assessing both the quality of the worlds they generate and their consistency and responsiveness under exploration, interaction, and modification. We introduce HappyWorld-Bench, a comprehensive benchmark that evaluates whether generated worlds remain reliable as agents interact with them...
Zhi-Qi Bai, Ju-Nai Cai, Yi-Xin Chen et al.· 0 citations
Multimodal understanding models that can jointly judge text-to-image (T2I), text-to-video (T2V) and text-to-speech (TTS) generation are increasingly used as"OmniJudges"for evaluation and automatic annotation. How reliably they understand what they score remains unclear, since existing benchmarks and training data tend...
Guang Hu, Ziyue Jiang, Wei Qiao et al.· 0 citations
Reward models play an essential role in aligning visual generative models, yet most existing visual reward models use a single scalar score or rely on fixed criteria that cannot adapt to different instructions. This limits both interpretability and task sensitivity, especially for text-to-image generation and instructi...
Zi-Jian Kan, Wei Wang, Long Luo et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.