The rapid advancement of AIGC video generation calls for evaluation frameworks that move beyond technical fidelity and incorporate human-centered aesthetic assessment. Existing benchmarks often overlook fine-grained perceptual qualities such as visual aesthetics, artistic style, and human preference. To address this li...
Long-Teng Jiang, Dan-Dan Zheng, Qian-Qian Qiao et al.· Proceedings of the Thirty-Fi...· 0 citations
PMOPD ( project-based Multi-Teacher On-Policy Distillation), which constructs subspace memories from the cumulative parameter displacements of different tasks and projects both gradients and optimizer updates to remove components that interfere with protected task directions, establishes geometry-aware optimization as...
You Liu, Ruo-Bing Zheng, Bo-Yuan Tong et al.· 0 citations
Making prediction with multimodal data is widely used in diverse scenarios. Existing multimodal fusion models, once deployed, can only handle predefined modalities (e.g., vision, text and audio) and single task, making it difficult to quickly adapt to new downstream applications. Therefore, a natural yet aggressive que...
Hui-Zi Cui, Zong-Bo Han, Chen Ding et al.· 0 citations
Video understanding is a fundamental task for evaluating the capabilities of multimodal large language models (MLLMs). However, existing leading models have already achieved approximately 90% accuracy on the Video-MME leaderboard, suggesting that conventional single-turn video understanding tasks are becoming increasin...
Fan Zhang, Guang-Ming Yao, Jin-Yang Wu et al.· 1 citation
SkySense-VITA is introduced, a unified in-context segmentation model, which synergistically processes both optical and Synthetic Aperture Radar (SAR) imagery using VI sual, T extu A l, or fused prompts, and is proposed to de-couple multi-modal prompt fusion and prediction process.
This work forms object-centric spatial intelligence as a unified problem connecting pose perception, spatial reasoning, pose-conditioned generation, and object-centric novel view synthesis and proposes a viewpoint-based orientation abstraction to make pose usable by multimodal large language models.
Mi-Ning Tan, Yinuo Wang, Ziqi Zhou et al.· 0 citations
Compass is presented, the first unified multimodal framework that grounds composition-intent control in a single system spanning both composition perception and composition-guided generation, with a shared expert token $\tau_c$ as the central intent anchor.
Ziqi Zhou, Weize Quan, Mi-Ning Tan et al.· arXiv.org· 0 citations
A novel framework that leverages multiple individual photographs to generate identitycustomized video, and a comprehensive high-quality dataset of 20,000 videos, thereby establishing a crucial resource to advance future research in multi-ID video generation.
Xinyang Song, Li-Bin Wang, Jianxin Sun et al.· IEEE transactions on multime...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.