With the rapid advancement of text-to-image (T2I) generation, robust evaluation becomes critical yet challenging, as traditional metrics fail to capture fine-grained alignment and generative artifacts. While large multimodal models (LMMs) are increasingly adopted as evaluators, existing benchmarks typically study seman...
Yu Zhao, Jia-Rui Wang, Hui-Yu Duan et al.· 0 citations
This work proposes LLaVA-Assessor, a unified data construction and model training system for LMM-based machine vision, and introduces a simple yet effective prompt disentanglement strategy to alleviate training-objective confusion in multi-task learning, thereby enabling stable and coherent joint training.
Zi-Heng Jia, Zi-Cheng Zhang, Jia-Ying Qian et al.· 0 citations
This work extends the previous dataset HVEval with pairwise preference annotations and proposes MoE-Rater, a Mixture-of-Experts (MoE)-inspired and multimodal large language model (MLLM)-based all-in-one method that supports multi-dimensional quality rating, multi-dimensional pairwise comparison, and category-specific q...
Si-Jing Wu, Yun-Hao Li, Hui-Yu Duan et al.· IEEE transactions on circuit...· 4 citations
MIE-Bench is introduced, the first large-scale multiple image editing benchmark with fine-grained human preference annotations and MIEScore, a multimodal large language model (MLLM)-based evaluation model enhanced with skill optimization and multi-dimensional supervised fine-tuning, to provide human-aligned feedback fo...
Zi-Tong Xu, Huiyu Duan, Xinyu Zhang et al.· 0 citations
This model leverages different layers of the large language model (LLM) decoder, treating them as multiple detectors that perform synchronous distortion detection using multi-level features, which helps mitigate the ambiguous foreground-background separation commonly encountered in the S2A problem.
Zi-Heng Jia, Yingji Liang, Jia-Ying Qian et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.