OmniJudge or OmniBias? Diagnosing Multimodal Judges through Balanced, Decoupled Lenses
Multimodal understanding models that can jointly judge text-to-image (T2I), text-to-video (T2V) and text-to-speech (TTS) generation are increasingly used as"OmniJudges"for evaluation and automatic annotation. How reliably they understand what they score remains unclear, since existing benchmarks and training data tend...