Oct 2026· IEEE Transactions on Emerging Topics in Computational Intelligence· Vol 10, pp. 3586-3601· 0 citations· 66 references
Abstract
Conventional Blind Image Quality Assessment (BIQA) methods typically assess the entire image quality, which is suboptimal for tasks like autonomous driving that concern specific Task-Aligned Region (TAR). Moreover, we observe that advanced Multimodal Large Language Model (MLLM)-based BIQA models exhibit bias when evaluating small regions, leading to inaccurate perceptual judgments. To address these issues, we propose SageIQ (Scene-graph-guided Evaluation for Image Quality), a pipeline SageIQ-P for TAR localization, an approach consisting of an MLLM-based BIQA model SageIQ-M, and a dataset SageIQ-D. SageIQ-P is designed to automatically identify and evaluate TARs, with the advantages of being training-free and allowing plug-in integration of off-the-shelf BIQA models without retraining. It operates in three stages: scenegraphbased triplet construction, LLMdriven triplet analysis, and integration of weighted BIQA scores into a final assessment. Since SageIQ-P can produce small-sized TAR crops that may encounter small-region scoring bias in existing BIQA models, we propose SageIQ-M to alleviate this bias by injecting scale information through scale-aware images and size-prompted cues, achieving size awareness across both visual and textual modalities. In addition, we develop a fully automated approach to construct a region-level test set SageIQ-D, significantly reducing the human effort needed. Experimental results demonstrate that our methods achieve superior BIQA performance.
Region-level No-Reference Image Quality Assessment (NR-IQA) enables fine-grained quality assessment for user-specified regions, which is critical in applications, such as camera imaging, autonomous driving and image compression. While existing NR-IQA methods have achieved remarkable success in assessing full-image qual...
Ze-Wen Chen, Juan Wang, Wen Wang et al.· IEEE Transactions on Image P...· 0 citations
A novel Hierarchical Multi-Scale Cross-Attention Network that effectively captures both local distortion patterns and global semantic information for quality prediction and exhibits superior generalization capability compared to existing approaches is proposed.
Jiakuo Yan, Jun Zeng· International Conference on...· 0 citations
Multispectral pan-sharpening aims to fuse high-resolution panchromatic and low-resolution multispectral imagery. However, this process introduces spatial artifacts and spectral distortions. Assessing the quality of fused images remains a fundamental challenge due to the absence of full-resolution ground-truth data. Thi...
Igor Stępień, Mariusz Oszust· Remote Sensing· 0 citations
This work proposes LLaVA-Assessor, a unified data construction and model training system for LMM-based machine vision, and introduces a simple yet effective prompt disentanglement strategy to alleviate training-objective confusion in multi-task learning, thereby enabling stable and coherent joint training.
Zi-Heng Jia, Zi-Cheng Zhang, Jia-Ying Qian et al.· 0 citations
Recent text-to-image models have made substantial progress in photorealism, aesthetics, and text-image alignment. Yet visually appealing images can still violate real-world plausibility, exhibiting malformed object structures, impossible anatomy, physically implausible interactions, or inconsistent spatial relationship...
In recent years, text-to-image (T2I) generation models have made substantial progress, particularly in visual realism and the expression of prompt semantics. However, a key difficulty remains: how to evaluate generated results automatically in a way that is both comprehensive and interpretable, while still being practi...