The rapid advancement of AIGC video generation calls for evaluation frameworks that move beyond technical fidelity and incorporate human-centered aesthetic assessment. Existing benchmarks often overlook fine-grained perceptual qualities such as visual aesthetics, artistic style, and human preference. To address this li...
Long-Teng Jiang, Dan-Dan Zheng, Qian-Qian Qiao et al.· Proceedings of the Thirty-Fi...· 0 citations
This work forms object-centric spatial intelligence as a unified problem connecting pose perception, spatial reasoning, pose-conditioned generation, and object-centric novel view synthesis and proposes a viewpoint-based orientation abstraction to make pose usable by multimodal large language models.
Mi-Ning Tan, Yinuo Wang, Ziqi Zhou et al.· 0 citations
Compass is presented, the first unified multimodal framework that grounds composition-intent control in a single system spanning both composition perception and composition-guided generation, with a shared expert token $\tau_c$ as the central intent anchor.
Ziqi Zhou, Weize Quan, Mi-Ning Tan et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.