Multi-modal large language models (MLLMs) have demonstrated significant potential in image quality assessment (IQA) by bridging visual perception with descriptive evaluations. However, existing approaches mainly focus on holistic quality prediction, often functioning as black boxes that provide limited insight into whe...
Ao-Ting Zhang, Ming-Ze Gao, Dong-Bao Yang et al.· 0 citations
This work introduces Emotion Statement Judgement (ESJ), a statement-verification formulation that preserves the expressiveness of the input space while constraining outputs to discriminative judgements, and builds EmObserver, an emotion-oriented MLLM optimized on ESJ through an elaborate multi-stage recipe.
Daiqing Wu, Dong-Bao Yang, Jiashu Yao et al.· arXiv.org· 1 citation
To minimize knowledge interference during fusion, this work presents a gradient-based orthogonal refreshing strategy that projects gradient updates of new domains onto the orthogonal complement of the fused historical subspace, supporting continual adaptation without forgetting.
Aoting Zhang, Dongbao Yang, Chang Liu et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.