Multimodal understanding models that can jointly judge text-to-image (T2I), text-to-video (T2V) and text-to-speech (TTS) generation are increasingly used as"OmniJudges"for evaluation and automatic annotation. How reliably they understand what they score remains unclear, since existing benchmarks and training data tend...
Guang Hu, Ziyue Jiang, Wei Qiao et al.· 0 citations
Reward models play an essential role in aligning visual generative models, yet most existing visual reward models use a single scalar score or rely on fixed criteria that cannot adapt to different instructions. This limits both interpretability and task sensitivity, especially for text-to-image generation and instructi...
Zi-Jian Kan, Wei Wang, Long Luo et al.· 0 citations
MoSaiC, a novel Motion-Saliency Complementary masked modeling framework for self-supervised point cloud video representation learning, couples three components: Curriculum Motion-Saliency Masking (CMSM), which guides the masking process toward motion-salient tokens under a curriculum schedule; Normal-Flow Motion (NFM)...
Wei Wang, Yi-Ding Sun, Yuying Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.